Papers with cross-modal task
Conditioned Masked Language and Image Modeling for Image-Text Dense Retrieval (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale two-stream pre-trained models like CLIP have achieved tremendous success in image-text retrieval. |
| Approach: | They propose a cross-modal framework for image-text retrieval using two-stream pre-trained models . they embed images and texts into instance representations with two separate encoders . experimental results on MSCOCO and Flickr30k reveal the effectiveness of their framework . |
| Outcome: | The proposed framework improves image-text retrieval performance on two popular cross-modal retrieval benchmarks. |
Extending CLIP’s Image-Text Alignment to Referring Image Segmentation (2024.naacl-long)
Copied to clipboard
| Challenge: | Referring Image Segmentation (RIS) is a cross-modal task that aims to segment an instance described by a natural language expression. |
| Approach: | They propose a framework that leverages the cross-modal nature of CLIP for RIS by leveraging image-text alignment knowledge in CLIP's image-embedding space. |
| Outcome: | The proposed framework outperforms CLIP-based methods on all three major RIS benchmarks and outperformed previous CLIP methods. |
On the Language Encoder of Contrastive Cross-modal Models (2024.findings-acl)
Copied to clipboard
Mengjie Zhao, Junya Ono, Zhi Zhong, Chieh-Hsin Lai, Yuhta Takida, Naoki Murata, Wei-Hsiang Liao, Takashi Shibuya, Hiromi Wakaki, Yuki Mitsufuji
| Challenge: | Pretrained audio-language models such as AudioCLIP and AudioCLAP have shown promising results on vision-language (VL) tasks. |
| Approach: | They extensively evaluate how unsupervised and supervised sentence embedding training affect language encoder quality and cross-modal task performance. |
| Outcome: | The proposed model improves on visual-language (VL) and audio-language tasks when the amount of training data is large. |
Representation Purification for End-to-End Speech Translation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing approaches to enhance speech translation focus on enhancing knowledge transfer . factors in speech that are not relevant to translation content, such as timbre and rhythm, often limit the efficiency of knowledge transfer. |
| Approach: | They propose a framework that excludes content-agnostic perturbations from speech representations to mitigate their negative impact on ST. |
| Outcome: | The proposed framework significantly improves translation performance across all translation directions in three settings and achieves preeminent performance under a *transcript-free* setting. |
CMOT: Cross-modal Mixup via Optimal Transport for Speech Translation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to translate speech signals into text are limited by the modality gap between speech and text. |
| Approach: | They propose Cross-modal Mixup via Optimal Transport to overcome the modality gap between speech and text by finding alignment between modalities. |
| Outcome: | Experiments on the MuST-C ST benchmark show that CMOT achieves an average BLEU of 30.0 in 8 translation directions, outperforming previous methods. |
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for visual storytelling ignore latent topic information. |
| Approach: | They propose a topic-aware reinforcement network for VIsual StoryTelling that takes topic information into account to generate a coherent story. |
| Outcome: | The proposed method outperforms most of the competing models across multiple evaluation metrics. |